Entrepreneurship

The Rise of Vals and the Urgent Need for Rigorous AI Model Benchmarking

In the rapidly expanding landscape of artificial intelligence, the standard practice of benchmarking has evolved from a technical necessity into a high-stakes marketing tool. AI developers frequently utilize performance metrics to validate their models’ capabilities, using favorable results to signal superiority over competitors. However, as the industry matures, a critical flaw has emerged: the reliance on antiquated, publicly available testing frameworks. These legacy systems, often ill-equipped to measure the nuanced performance of modern large language models (LLMs), have become susceptible to "gaming." When models are trained on the very data contained in public benchmarks, the results cease to be an objective measurement of intelligence and instead become a reflection of data memorization.

Enter Vals, a San Francisco-based startup founded in 2024 with the explicit mission to revolutionize how we measure artificial intelligence. In less than two years, the company has transitioned from a nascent idea to a cornerstone of AI evaluation, recently securing $40 million in a Series A funding round led by Andreessen Horowitz. This financial milestone follows a successful seed round backed by prominent firms including 8VC and Bloomberg Beta, signaling strong investor confidence in the necessity of independent, robust AI auditing.

The Architect of a New Evaluation Standard

The genesis of Vals can be traced to the observations of its 25-year-old co-founder, Rayan Krishnan. A Stanford alumnus with professional experience at Microsoft’s AI labs and an internship at Palantir, Krishnan witnessed firsthand the widening gap between the rapid deployment of frontier AI models and the static, academic benchmarks used to assess them.

“We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks were not keeping up with that frontier advance,” Krishnan explains. His perspective is grounded in the belief that as AI integration permeates critical societal sectors—from healthcare to legal services—benchmarks must evolve beyond abstract testing. The current reliance on static datasets, he argues, does little to verify if a model can perform complex, real-world tasks with human-level accuracy.

Operating out of a historic brewery building on Folsom Street in San Francisco—a structure that now serves as a hub for emerging tech ventures—Vals has adopted an approach that prioritizes private, dynamic evaluation. By keeping its test materials confidential, the company prevents the common practice of "exam cheating," where developers inadvertently or intentionally train models on benchmark content.

Moving Beyond General Knowledge

Traditional benchmarking often centers on general intelligence, tasking models with answering trivia questions or solving standardized exams like the bar exam or medical boards. While these metrics provide a baseline, they do not account for the operational reliability required in professional environments. Vals shifts the focus toward domain-specific competence, testing models on their ability to execute complex workflows in industries such as finance, software engineering, and legal analysis.

The company’s methodology involves evaluating not just the potential for positive outcomes, but the inherent risks associated with model autonomy. This includes testing for "negative implications" should models be deployed without sufficient guardrails. The scope of their evaluation is expansive, covering sensitive and high-stakes areas such as cybersecurity, biosecurity, mental health, and the application of international law, including the Geneva Convention. By stress-testing these models against adversarial scenarios, Vals aims to provide a comprehensive risk profile that traditional benchmarks fail to capture.

The Economics of Verification

The business model of Vals relies on a counterintuitive premise: companies paying to have their own products scrutinized for potential failures. However, this model mirrors the professional certification processes found in other high-stakes industries. Much like a student pays the College Board to take the SAT, or an engineer pays for professional certification, AI developers are increasingly recognizing the value of third-party validation to identify vulnerabilities and refine performance.

Data indicates that this market for truth is growing rapidly. Vals has reported an eightfold increase in revenue year-over-year, and its internal headcount has tripled from eight to 25 employees in less than 12 months. With plans to hire an additional 10 to 15 professionals and move into a larger facility, the company is positioning itself as a vital infrastructure provider for the AI era. Furthermore, the launch of a dedicated program for federal agencies underscores the strategic importance of independent evaluation for national security and public policy.

Implications for the Future of AI Transparency

The transition of AI companies from private entities to publicly traded corporations is expected to fundamentally change the requirements for model documentation. With major players like Anthropic rumored to be exploring public offerings and the potential for others like OpenAI to follow suit, the demand for standardized, reliable auditing will only intensify.

Krishnan suggests that the benchmarks and evaluation reports produced by firms like Vals will eventually become a staple of public financial filings. As AI models become central to the global economy, investors and regulators will require more than marketing claims to assess the long-term viability and safety of these systems. “I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they’re going to make,” Krishnan notes.

A Critical Shift in Industry Standards

The rise of Vals represents a broader shift in the tech industry toward accountability. For years, the lack of standardized evaluation allowed for a "Wild West" environment where capabilities were often inflated. By providing a third-party, adversarial, and confidential evaluation service, Vals is helping to build the "trust infrastructure" necessary for enterprise adoption.

However, the challenge remains significant. As models become more recursive—capable of self-improvement and complex reasoning—the benchmarks must become equally sophisticated. The work being done by Vals in areas like recursive self-improvement and biosecurity is a reflection of the industry’s growing awareness that the risks associated with frontier AI are not merely theoretical.

For the enterprise sector, the implications are clear: the future of AI procurement will be determined by verified performance data rather than marketing rhetoric. By establishing a rigorous, industry-agnostic standard for testing, Vals is attempting to ensure that as AI models are integrated into the fabric of society, they do so with a verified level of safety, reliability, and human-comparable quality. As the company continues to scale, its ability to maintain independence and technical rigor will be the primary metric by which its own success—and the success of the broader AI benchmarking industry—is measured.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button